<!DOCTYPE html>
<html class="client-nojs vector-feature-night-mode-disabled vector-feature-language-in-header-enabled vector-feature-language-in-main-page-header-disabled vector-feature-page-tools-pinned-disabled vector-feature-toc-pinned-clientpref-1 vector-feature-main-menu-pinned-disabled vector-feature-limited-width-clientpref-1 vector-feature-limited-width-content-enabled vector-feature-custom-font-size-clientpref-1 vector-feature-appearance-pinned-clientpref-1 vector-sticky-header-enabled" lang="en" dir="ltr"><head>
<meta charset="UTF-8">
<title>Predication (computer architecture)</title>
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<link rel="canonical" href="https://en.wikipedia.org/wiki/Predication_(computer_architecture)"> <link href="./mw/ext.cite.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/ext.pygments.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.icons.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.search.codex.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/skins.vector.styles.css" rel="stylesheet" type="text/css">
<link href="./mw/user.styles.css" rel="stylesheet" type="text/css">
<meta name="ResourceLoaderDynamicStyles" content="">
<link rel="stylesheet" type="text/css" href="./mw/site.styles.css">
<link rel="stylesheet" type="text/css" href="./mw/noscript.css">
<link rel="stylesheet" type="text/css" href="./footer.css">
<link rel="stylesheet" type="text/css" href="./vector-2022.css">
</head>
<body class="skin--responsive skin-vector skin-vector-search-vue mediawiki ltr sitedir-ltr mw-hide-empty-elt ns-0 ns-subject page-Predication_computer_architecture rootpage-Predication_computer_architecture skin-vector-2022 action-view">
<div class="mw-page-container">
<div class="mw-page-container-inner">
<div class="mw-content-container">
<main id="content" class="mw-body">
<header class="mw-body-header vector-page-titlebar">
<h1 id="firstHeading" class="firstHeading mw-first-heading">
<span id="openzim-page-title" class="mw-page-title-main"><span class="mw-page-title-main">Predication (computer architecture)</span></span>
</h1>
</header>
<a id="top"></a>
<div id="bodyContent" class="vector-body ve-init-mw-desktopArticleTarget-targetContainer" aria-labelledby="firstHeading" data-mw-ve-target-container="">
<div id="mw-content-text" class="mw-body-content mw-content-ltr" lang="en" dir="ltr"><div class="mw-content-ltr mw-parser-output" lang="en" dir="ltr">
<style data-mw-deduplicate="TemplateStyles:r1305420660">
/* start https://en.wikipedia.org/ */
.mw-parser-output .ambox{border:1px solid #a2a9b1;box-sizing:border-box}.mw-parser-output .ambox+link+.ambox,.mw-parser-output .ambox+link+style+.ambox,.mw-parser-output .ambox+link+link+.ambox,.mw-parser-output .ambox+.mw-empty-elt+link+.ambox,.mw-parser-output .ambox+.mw-empty-elt+link+style+.ambox,.mw-parser-output .ambox+.mw-empty-elt+link+link+.ambox{margin-top:-1px}html body.mediawiki .mw-parser-output .ambox.mbox-small-left{margin:4px 1em 4px 0;overflow:hidden;width:238px;border-collapse:collapse;font-size:88%;line-height:1.25em}html body.mediawiki .mw-parser-output .metadata.ambox{border-left:10px solid #36c!important;background-color:var(--background-color-neutral-subtle,#fbfbfb)!important}html body.mediawiki .mw-parser-output .metadata.ambox-speedy{border-left:10px solid #b32424!important;background-color:var(--background-color-destructive-subtle,#fee7e6)!important}html body.mediawiki .mw-parser-output .metadata.ambox-delete{border-left:10px solid #b32424!important}html body.mediawiki .mw-parser-output .metadata.ambox-content{border-left:10px solid #f28500!important}html body.mediawiki .mw-parser-output .metadata.ambox-style{border-left:10px solid #fc3!important}html body.mediawiki .mw-parser-output .metadata.ambox-move{border-left:10px solid #9932cc!important}html body.mediawiki .mw-parser-output .metadata.ambox-protection{border-left:10px solid #a2a9b1!important}.mw-parser-output .ambox.mbox-text{border:none;padding:0.25em 0.5em;width:100%}.mw-parser-output .ambox .mbox-image{border:none;padding:2px 0 2px 0.5em;text-align:center}.mw-parser-output .ambox .mbox-imageright{border:none;padding:2px 0.5em 2px 0;text-align:center}.mw-parser-output .ambox .mbox-empty-cell{border:none;padding:0;width:1px}.mw-parser-output .ambox .mbox-image-div{width:52px}@media(min-width:720px){.mw-parser-output .ambox{margin:0 10%}}@media print{body.ns-0 .mw-parser-output .ambox{display:none!important}}
/* end https://en.wikipedia.org/ */
</style>
<style data-mw-deduplicate="TemplateStyles:r1236090951">
/* start https://en.wikipedia.org/ */
.mw-parser-output .hatnote{font-style:italic}.mw-parser-output div.hatnote{padding-left:1.6em;margin-bottom:0.5em}.mw-parser-output .hatnote i{font-style:normal}.mw-parser-output .hatnote+link+.hatnote{margin-top:-0.5em}@media print{body.ns-0 .mw-parser-output .hatnote{display:none!important}}
/* end https://en.wikipedia.org/ */
</style><div role="note" class="hatnote navigation-not-searchable">Not to be confused with <a href="Branch_prediction" class="mw-redirect" title="Branch prediction">branch prediction</a>.</div>
<p>In <a href="Computer_architecture" title="Computer architecture">computer architecture</a>, <b>predication</b> is a feature that provides an alternative to <a href="Conditional_(computer_programming)" title="Conditional (computer programming)">conditional</a> transfer of <a href="Control_flow" title="Control flow">control</a>, as implemented by conditional <a href="Branch_(computer_science)" title="Branch (computer science)">branch</a> machine <a href="Instruction_(computer_science)" class="mw-redirect" title="Instruction (computer science)">instructions</a>. Predication works by having conditional (<i>predicated</i>) non-branch instructions associated with a <i>predicate</i>, a <a href="Boolean_data_type" title="Boolean data type">Boolean value</a> used by the instruction to control whether the instruction is allowed to modify the architectural state or not. If the predicate specified in the instruction is true, the instruction modifies the architectural state; otherwise, the architectural state is unchanged. For example, a predicated move instruction (a conditional move) will only modify the destination if the predicate is true. Thus, instead of using a conditional branch to select an instruction or a sequence of instructions to <a href="Execution_(computing)" title="Execution (computing)">execute</a> based on the predicate that controls whether the branch occurs, the instructions to be executed are associated with that predicate, so that they will be executed, or not executed, based on whether that predicate is true or false.<sup id="cite_ref-rvinyard_1-0" class="reference"><a href="#cite_note-rvinyard-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p><p><a href="Vector_processors" class="mw-redirect" title="Vector processors">Vector processors</a>, some <a href="SIMD" class="mw-redirect" title="SIMD">SIMD</a> ISAs (such as <a href="AVX2" class="mw-redirect" title="AVX2">AVX2</a> and <a href="AVX-512" title="AVX-512">AVX-512</a>) and <a href="GPU" class="mw-redirect" title="GPU">GPUs</a> in general make heavy use of predication, applying one bit of a conditional <i>mask vector</i> to the corresponding elements in the vector registers being processed, whereas scalar predication in scalar instruction sets only need the one predicate bit. Where predicate masks become particularly powerful in <a href="Vector_processing" class="mw-redirect" title="Vector processing">vector processing</a> is if an <i>array</i> of <a href="Condition_code_register" class="mw-redirect" title="Condition code register">condition codes</a>, one per vector element, may feed back into predicate masks that are then applied to subsequent vector instructions.
</p>
<meta property="mw:PageProp/toc">
<div class="mw-heading mw-heading2"><h2 id="Overview">Overview</h2></div>
<p>Most <a href="Computer_program" title="Computer program">computer programs</a> contain <a href="Conditional_(computer_programming)" title="Conditional (computer programming)">conditional</a> code, which will be executed only under specific conditions depending on factors that cannot be determined beforehand, for example depending on user input. As the majority of <a href="Central_processing_unit" title="Central processing unit">processors</a> simply execute the next <a href="Instruction_(computer_science)" class="mw-redirect" title="Instruction (computer science)">instruction</a> in a sequence, the traditional solution is to insert <i>branch</i> instructions that allow a program to conditionally branch to a different section of code, thus changing the next step in the sequence. This was sufficient until designers began improving performance by implementing <a href="Instruction_pipelining" title="Instruction pipelining">instruction pipelining</a>, a method which is slowed down by branches. For a more thorough description of the problems which arose, and a popular solution, see <a href="Branch_predictor" title="Branch predictor">branch predictor</a>.
</p><p>Luckily, one of the more common patterns of code that normally relies on branching has a more elegant solution. Consider the following <a href="Pseudocode" title="Pseudocode">pseudocode</a>:<sup id="cite_ref-rvinyard_1-1" class="reference"><a href="#cite_note-rvinyard-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-highlight mw-highlight-lang-c mw-content-ltr" dir="ltr"><pre><span class="k">if</span><span class="w"> </span><span class="n">condition</span>
<span class="w"> </span><span class="p">{</span><span class="n">do_something</span><span class="p">}</span>
<span class="k">else</span>
<span class="w"> </span><span class="p">{</span><span class="n">do_something_else</span><span class="p">}</span>
</pre></div>
<p>On a system that uses conditional branching, this might translate to <a href="Machine_language" class="mw-redirect" title="Machine language">machine instructions</a> looking similar to:<sup id="cite_ref-rvinyard_1-2" class="reference"><a href="#cite_note-rvinyard-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-highlight mw-highlight-lang-c mw-content-ltr" dir="ltr"><pre><span class="w"> </span><span class="n">branch_if_condition_to</span><span class="w"> </span><span class="n">label1</span>
<span class="w"> </span><span class="n">do_something_else</span>
<span class="w"> </span><span class="n">branch_always_to</span><span class="w"> </span><span class="n">label2</span>
<span class="nl">label1</span><span class="p">:</span>
<span class="w"> </span><span class="n">do_something</span>
<span class="nl">label2</span><span class="p">:</span>
<span class="w"> </span><span class="p">...</span>
</pre></div>
<p>With predication, all possible branch paths are coded inline, but some instructions execute while others do not. The basic idea is that each instruction is associated with a predicate (the word here used similarly to its usage in <a href="Predicate_logic" class="mw-redirect" title="Predicate logic">predicate logic</a>) and that the instruction will only be executed if the predicate is true. The machine code for the above example using predication might look something like this:<sup id="cite_ref-rvinyard_1-3" class="reference"><a href="#cite_note-rvinyard-1"><span class="cite-bracket">[</span>1<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-highlight mw-highlight-lang-c mw-content-ltr" dir="ltr"><pre><span class="p">(</span><span class="n">condition</span><span class="p">)</span><span class="w"> </span><span class="n">do_something</span>
<span class="p">(</span><span class="n">not</span><span class="w"> </span><span class="n">condition</span><span class="p">)</span><span class="w"> </span><span class="n">do_something_else</span>
</pre></div>
<p>Besides eliminating branches, less code is needed in total, provided the architecture provides predicated instructions. While this does not guarantee faster execution in general, it will if the <code>do_something</code> and <code>do_something_else</code> blocks of code are short enough.
</p><p>Predication's simplest form is <i>partial predication</i>, where the architecture has <i>conditional move</i> or <i>conditional select</i> instructions. Conditional move instructions write the contents of one register over another only if the predicate's value is true, whereas conditional select instructions choose which of two registers has its contents written to a third based on the predicate's value. A more generalized and capable form is <i>full predication</i>. Full predication has a set of predicate registers for storing predicates (which allows multiple nested or sequential branches to be simultaneously eliminated) and most instructions in the architecture have an (optional) register specifier field to specify which predicate register supplies the predicate.<sup id="cite_ref-2" class="reference"><a href="#cite_note-2"><span class="cite-bracket">[</span>2<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="Advantages">Advantages</h2></div>
<p>The main purpose of predication is to avoid jumps over very small sections of program code, increasing the effectiveness of <a href="Pipeline_(computing)" title="Pipeline (computing)">pipelined</a> execution and avoiding problems with the <a href="CPU_cache" title="CPU cache">cache</a>. It also has a number of more subtle benefits:
</p>
<ul><li>Functions that are traditionally computed using simple arithmetic and <a href="Bitwise_operation" title="Bitwise operation">bitwise operations</a> may be quicker to compute using predicated instructions.</li>
<li>Predicated instructions with different predicates can be mixed with each other and with unconditional code, allowing better <a href="Instruction_scheduling" title="Instruction scheduling">instruction scheduling</a> and so even better performance.</li>
<li>Elimination of unnecessary branch instructions can make the execution of necessary branches, such as those that make up loops, faster by lessening the load on <a href="Branch_prediction" class="mw-redirect" title="Branch prediction">branch prediction</a> mechanisms.</li>
<li>Elimination of the cost of a branch misprediction which can be high on deeply pipelined architectures.</li>
<li>Instruction sets that have comprehensive <a href="Condition_code_register" class="mw-redirect" title="Condition code register">Condition Codes</a> generated by instructions may reduce code size further by directly using the Condition Registers in or as predication.</li></ul>
<div class="mw-heading mw-heading2"><h2 id="Disadvantages">Disadvantages</h2></div>
<p>Predication's primary drawback is in increased encoding space. In typical implementations, every instruction reserves a bitfield for the predicate specifying under what conditions that instruction should have an effect. When available memory is limited, as on <a href="Embedded_system" title="Embedded system">embedded devices</a>, this space cost can be prohibitive. However, some architectures such as <a href="Thumb-2" class="mw-redirect" title="Thumb-2">Thumb-2</a> are able to avoid this issue (see below). Other detriments are the following:<sup id="cite_ref-Fisher04_3-0" class="reference"><a href="#cite_note-Fisher04-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup>
</p>
<ul><li>Predication complicates the hardware by adding levels of <a href="Control_unit" title="Control unit">logic</a> to critical <a href="Datapath" title="Datapath">paths</a> and potentially degrades clock speed.</li>
<li>A predicated block includes cycles for all operations, so shorter <a href="Control-flow_graph" title="Control-flow graph">paths</a> may take longer and be penalized.</li>
<li>An extra register read is required. A non-predicated ADD would read two registers from a register file, where a Predicated ADD would need to also read the predicate register file. This increases Hazards in <a href="Out-of-order_execution" title="Out-of-order execution">Out-of-order execution</a>.</li>
<li>Predication is not usually speculated and causes a longer dependency chain. For ordered data this translates to a performance loss compared to a predictable branch.<sup id="cite_ref-4" class="reference"><a href="#cite_note-4"><span class="cite-bracket">[</span>4<span class="cite-bracket">]</span></a></sup></li></ul>
<p>Predication is most effective when paths are balanced or when the longest path is the most frequently executed,<sup id="cite_ref-Fisher04_3-1" class="reference"><a href="#cite_note-Fisher04-3"><span class="cite-bracket">[</span>3<span class="cite-bracket">]</span></a></sup> but determining such a path is very difficult at compile time, even in the presence of <a href="Profiling_(computer_programming)" title="Profiling (computer programming)">profiling information</a>.
</p>
<div class="mw-heading mw-heading2"><h2 id="History">History</h2></div>
<p>Predicated instructions were popular in European computer designs of the 1950s, including the <a href="Mail%C3%BCfterl" title="Mailüfterl">Mailüfterl</a> (1955), the <a href="Z22_(computer)" title="Z22 (computer)">Zuse Z22</a> (1955), the <a href="ZEBRA_(computer)" title="ZEBRA (computer)">ZEBRA</a> (1958), and the <a href="Electrologica_X1" title="Electrologica X1">Electrologica X1</a> (1958). The <a href="ACS-1" class="mw-redirect" title="ACS-1">IBM ACS-1</a> design of 1967 allocated a "skip" bit in its instruction formats, and the CDC Flexible Processor in 1976 allocated three conditional execution bits in its microinstruction formats.
</p><p><a href="Hewlett-Packard" title="Hewlett-Packard">Hewlett-Packard</a>'s <a href="PA-RISC" title="PA-RISC">PA-RISC</a> architecture (1986) had a feature called <i>nullification</i>, which allowed most instructions to be predicated by the previous instruction. <a href="IBM" title="IBM">IBM</a>'s <a href="IBM_POWER_instruction_set_architecture" class="mw-redirect" title="IBM POWER instruction set architecture">POWER architecture</a> (1990) featured conditional move instructions. POWER's successor, <a href="PowerPC" title="PowerPC">PowerPC</a> (1993), dropped these instructions. <a href="Digital_Equipment_Corporation" title="Digital Equipment Corporation">Digital Equipment Corporation</a>'s <a href="DEC_Alpha" title="DEC Alpha">Alpha</a> architecture (1992) also featured conditional move instructions. <a href="MIPS_architecture" title="MIPS architecture">MIPS</a> gained conditional move instructions in 1994 with the MIPS IV version; and <a href="SPARC" title="SPARC">SPARC</a> was extended in Version 9 (1994) with conditional move instructions for both integer and floating-point registers.
</p><p>In the <a href="Hewlett-Packard" title="Hewlett-Packard">Hewlett-Packard</a>/<a href="Intel" title="Intel">Intel</a> <a href="IA-64" title="IA-64">IA-64</a> architecture, most instructions are predicated. The predicates are stored in 64 special-purpose predicate <a href="Processor_register" title="Processor register">registers</a>; and one of the predicate registers is always true so that <i>unpredicated</i> instructions are simply instructions predicated with the value true. The use of predication is essential in IA-64's implementation of <a href="Software_pipelining" title="Software pipelining">software pipelining</a> because it avoids the need for writing separated code for prologs and epilogs.
</p><p>In the <a href="X86" title="X86">x86</a> architecture, a family of conditional move instructions (<code>CMOV</code> and <code>FCMOV</code>) were added to the architecture by the <a href="Intel" title="Intel">Intel</a> <a href="Pentium_Pro" title="Pentium Pro">Pentium Pro</a> (1995) processor. The <code>CMOV</code> instructions copied the contents of the source register to the destination register depending on a predicate supplied by the value of the flag register.
</p><p>In the <a href="ARM_architecture" class="mw-redirect" title="ARM architecture">ARM architecture</a>, the original 32-bit instruction set provides a feature called <i>conditional execution</i> that allows most instructions to be predicated by one of 13 predicates that are based on some combination of the four condition codes set by the previous instruction. ARM's <a href="ARM_architecture" class="mw-redirect" title="ARM architecture">Thumb</a> instruction set (1994) dropped conditional execution to reduce the size of instructions so they could fit in 16 bits, but its successor, <a href="ARM_architecture" class="mw-redirect" title="ARM architecture">Thumb-2</a> (2003) overcame this problem by using a special instruction which has no effect other than to supply predicates for the following four instructions. The 64-bit instruction set introduced in ARMv8-A (2011) replaced conditional execution with conditional selection instructions.
</p>
<div class="mw-heading mw-heading2"><h2 id="SIMD,_SIMT_and_vector_predication">SIMD, SIMT and vector predication</h2></div>
<div role="note" class="hatnote navigation-not-searchable">See also: <a href="Single_instruction%2C_multiple_threads" title="Single instruction, multiple threads">Single instruction, multiple threads</a></div>
<p>Some <a href="SIMD_within_a_register" class="mw-redirect" title="SIMD within a register">SIMD within a register</a> instruction sets, like AVX2, have the ability to use a logical <a href="Mask_(computing)" title="Mask (computing)">mask</a> to conditionally load/store values to memory, in a parallel form of the conditional move. They may also apply individual mask bits to individual arithmetic units executing a parallel operation. One predicate mask bit is available for each sub-word of the SWAR register or each load/store value.
This form of multi-bit predication is also used in <a href="Vector_processors" class="mw-redirect" title="Vector processors">vector processors</a> at the element level (synonymous with SWAR sub-words):
</p>
<div class="mw-highlight mw-highlight-lang-c mw-content-ltr" dir="ltr"><pre><span class="k">for</span><span class="w"> </span><span class="n">each</span><span class="w"> </span><span class="p">(</span><span class="n">sub</span><span class="o">-</span><span class="n">word</span><span class="w"> </span><span class="n">i</span><span class="p">)</span><span class="w"> </span><span class="n">of</span><span class="w"> </span><span class="n">SWAR</span><span class="w"> </span><span class="p">(</span><span class="n">or</span><span class="w"> </span><span class="n">Vector</span><span class="p">)</span><span class="w"> </span><span class="k">register</span>
<span class="w"> </span><span class="p">(</span><span class="n">condition</span><span class="o">-</span><span class="n">maskbit</span><span class="w"> </span><span class="n">i</span><span class="p">)</span><span class="w"> </span><span class="n">do_something</span><span class="p">(</span><span class="n">sub</span><span class="o">-</span><span class="n">word</span><span class="w"> </span><span class="n">i</span><span class="p">)</span>
<span class="w"> </span><span class="p">(</span><span class="n">not</span><span class="w"> </span><span class="n">condition</span><span class="o">-</span><span class="n">maskbit</span><span class="w"> </span><span class="n">i</span><span class="p">)</span><span class="w"> </span><span class="n">do_something_else</span><span class="p">(</span><span class="n">sub</span><span class="o">-</span><span class="n">word</span><span class="w"> </span><span class="n">i</span><span class="p">)</span>
</pre></div>
<p>Masking is an integral part of <a href="Flynn's_taxonomy" title="Flynn's taxonomy">Array Processors</a> such as the <a href="ILLIAC_IV" title="ILLIAC IV">ILLIAC IV</a>. Array Processors are known today as <a href="Single_instruction%2C_multiple_threads" title="Single instruction, multiple threads">single instruction, multiple threads</a> (SIMT), and a predicate bit <i>per PE</i> used to activate or de-activate each Processing Element. When the PE has no <a href="SIMD_within_a_register" class="mw-redirect" title="SIMD within a register">SIMD within a register</a> instructions, each PE may be individually Predicated:
</p>
<div class="mw-highlight mw-highlight-lang-c mw-content-ltr" dir="ltr"><pre><span class="k">for</span><span class="w"> </span><span class="n">each</span><span class="w"> </span><span class="p">(</span><span class="n">PE</span><span class="w"> </span><span class="n">j</span><span class="p">)</span><span class="w"> </span><span class="c1">// of non-SWAR synchronously-concurrent array</span>
<span class="w"> </span><span class="p">(</span><span class="n">active</span><span class="o">-</span><span class="n">maskbit</span><span class="w"> </span><span class="n">j</span><span class="p">)</span><span class="w"> </span><span class="n">broadcast_scalar_instruction_to</span><span class="p">(</span><span class="n">PE</span><span class="w"> </span><span class="n">j</span><span class="p">)</span>
</pre></div>
<p>Modern SIMT <a href="GPUs" class="mw-redirect" title="GPUs">GPUs</a> use (or used, but ILLIAC IV documentation termed it <a href="ILLIAC_IV#Branches" title="ILLIAC IV">"branching"</a>) predication to enable/disable individual Processing Elements <i>and</i>, separately and furthermore, to <i>also</i> mask-out sub-words within any given PE's SWAR ALU.
</p>
<div class="mw-highlight mw-highlight-lang-c mw-content-ltr" dir="ltr"><pre><span class="k">for</span><span class="w"> </span><span class="n">each</span><span class="w"> </span><span class="p">(</span><span class="n">PE</span><span class="w"> </span><span class="n">j</span><span class="p">)</span><span class="w"> </span><span class="n">of</span><span class="w"> </span><span class="n">SIMT</span><span class="w"> </span><span class="n">synchronously</span><span class="o">-</span><span class="n">concurrent</span><span class="w"> </span><span class="n">array</span>
<span class="w"> </span><span class="p">(</span><span class="n">active</span><span class="o">-</span><span class="n">maskbit</span><span class="w"> </span><span class="n">j</span><span class="p">)</span><span class="w"> </span><span class="p">{</span><span class="w"> </span><span class="c1">// broadcast only to active SWAR PEs</span>
<span class="w"> </span><span class="k">for</span><span class="w"> </span><span class="n">each</span><span class="w"> </span><span class="p">(</span><span class="n">sub</span><span class="o">-</span><span class="n">word</span><span class="w"> </span><span class="n">i</span><span class="p">)</span><span class="w"> </span><span class="n">of</span><span class="w"> </span><span class="n">SWAR</span><span class="w"> </span><span class="k">register</span><span class="w"> </span><span class="n">in</span><span class="w"> </span><span class="p">(</span><span class="n">PE</span><span class="w"> </span><span class="n">j</span><span class="p">)</span>
<span class="w"> </span><span class="p">(</span><span class="n">condition</span><span class="o">-</span><span class="n">maskbit</span><span class="w"> </span><span class="n">i</span><span class="p">)</span><span class="w"> </span><span class="n">do_something</span><span class="p">(</span><span class="n">sub</span><span class="o">-</span><span class="n">word</span><span class="w"> </span><span class="n">i</span><span class="p">)</span>
<span class="w"> </span><span class="p">(</span><span class="n">not</span><span class="w"> </span><span class="n">condition</span><span class="o">-</span><span class="n">maskbit</span><span class="w"> </span><span class="n">i</span><span class="p">)</span><span class="w"> </span><span class="n">do_something_else</span><span class="p">(</span><span class="n">sub</span><span class="o">-</span><span class="n">word</span><span class="w"> </span><span class="n">i</span><span class="p">)</span>
<span class="w"> </span><span class="p">}</span>
</pre></div>
<p>All the techniques, advantages and disadvantages of single scalar predication apply just as well to the parallel processing case, where the issues associated with branching are made far more complex.<sup id="cite_ref-5" class="reference"><a href="#cite_note-5"><span class="cite-bracket">[</span>5<span class="cite-bracket">]</span></a></sup>
</p>
<div class="mw-heading mw-heading2"><h2 id="See_also">See also</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1184024115">
/* start https://en.wikipedia.org/ */
.mw-parser-output .div-col{margin-top:0.3em;column-width:30em}.mw-parser-output .div-col-small{font-size:90%}.mw-parser-output .div-col-rules{column-rule:1px solid #aaa}.mw-parser-output .div-col dl,.mw-parser-output .div-col ol,.mw-parser-output .div-col ul{margin-top:0}.mw-parser-output .div-col li,.mw-parser-output .div-col dd{page-break-inside:avoid;break-inside:avoid-column}
/* end https://en.wikipedia.org/ */
</style><div class="div-col" style="column-width: 20em;">
<ul><li><a href="Branch_predictor" title="Branch predictor">Branch predictor</a></li>
<li><a href="Control_flow" title="Control flow">Control flow</a></li>
<li><a href="Delay_slot" title="Delay slot">Delay slot</a></li>
<li><a href="Instruction-level_parallelism" title="Instruction-level parallelism">Instruction-level parallelism</a></li>
<li><a href="Optimizing_compiler" title="Optimizing compiler">Optimizing compiler</a></li>
<li><a href="Pipeline_stall" title="Pipeline stall">Pipeline stall</a></li>
<li><a href="Software_pipelining" title="Software pipelining">Software pipelining</a></li>
<li><a href="Speculative_execution" title="Speculative execution">Speculative execution</a></li>
<li><a href="Vector_processor" title="Vector processor">Vector processor</a></li>
<li><a href="Very_long_instruction_word" title="Very long instruction word">Very long instruction word</a></li></ul>
</div>
<div class="mw-heading mw-heading2"><h2 id="References">References</h2></div>
<style data-mw-deduplicate="TemplateStyles:r1239543626">
/* start https://en.wikipedia.org/ */
.mw-parser-output .reflist{margin-bottom:0.5em;list-style-type:decimal}@media screen{.mw-parser-output .reflist{font-size:90%}}.mw-parser-output .reflist .references{font-size:100%;margin-bottom:0;list-style-type:inherit}.mw-parser-output .reflist-columns-2{column-width:30em}.mw-parser-output .reflist-columns-3{column-width:25em}.mw-parser-output .reflist-columns{margin-top:0.3em}.mw-parser-output .reflist-columns ol{margin-top:0}.mw-parser-output .reflist-columns li{page-break-inside:avoid;break-inside:avoid-column}.mw-parser-output .reflist-upper-alpha{list-style-type:upper-alpha}.mw-parser-output .reflist-upper-roman{list-style-type:upper-roman}.mw-parser-output .reflist-lower-alpha{list-style-type:lower-alpha}.mw-parser-output .reflist-lower-greek{list-style-type:lower-greek}.mw-parser-output .reflist-lower-roman{list-style-type:lower-roman}
/* end https://en.wikipedia.org/ */
</style><div class="reflist">
<div class="mw-references-wrap"><ol class="references">
<li id="cite_note-rvinyard-1"><span class="mw-cite-backlink">^ <a href="#cite_ref-rvinyard_1-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-rvinyard_1-1"><sup><i><b>b</b></i></sup></a> <a href="#cite_ref-rvinyard_1-2"><sup><i><b>c</b></i></sup></a> <a href="#cite_ref-rvinyard_1-3"><sup><i><b>d</b></i></sup></a></span> <span class="reference-text"><style data-mw-deduplicate="TemplateStyles:r1238218222">
/* start https://en.wikipedia.org/ */
.mw-parser-output cite.citation{font-style:inherit;word-wrap:break-word}.mw-parser-output .citation q{quotes:"\"""\"""'""'"}.mw-parser-output .citation:target{background-color:rgba(0,127,255,0.133)}.mw-parser-output .id-lock-free.id-lock-free a{background:url("./mw/Lock-green.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-limited.id-lock-limited a,.mw-parser-output .id-lock-registration.id-lock-registration a{background:url("./mw/Lock-gray-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .id-lock-subscription.id-lock-subscription a{background:url("./mw/Lock-red-alt-2.svg")right 0.1em center/9px no-repeat}.mw-parser-output .cs1-ws-icon a{background:url("./mw/Wikisource-logo.svg")right 0.1em center/12px no-repeat}body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-free a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-limited a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-registration a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .id-lock-subscription a,body:not(.skin-timeless):not(.skin-minerva) .mw-parser-output .cs1-ws-icon a{background-size:contain;padding:0 1em 0 0}.mw-parser-output .cs1-code{color:inherit;background:inherit;border:none;padding:inherit}.mw-parser-output .cs1-hidden-error{display:none;color:var(--color-error,#d33)}.mw-parser-output .cs1-visible-error{color:var(--color-error,#d33)}.mw-parser-output .cs1-maint{display:none;color:#085;margin-left:0.3em}.mw-parser-output .cs1-kern-left{padding-left:0.2em}.mw-parser-output .cs1-kern-right{padding-right:0.2em}.mw-parser-output .citation .mw-selflink{font-weight:inherit}@media screen{.mw-parser-output .cs1-format{font-size:95%}html.skin-theme-clientpref-night .mw-parser-output .cs1-maint{color:#18911f}}@media screen and (prefers-color-scheme:dark){html.skin-theme-clientpref-os .mw-parser-output .cs1-maint{color:#18911f}}
/* end https://en.wikipedia.org/ */
</style><cite id="CITEREFRick_Vinyard2000" class="citation web cs1">Rick Vinyard (2000-04-26). <a rel="nofollow" class="external text" href="https://web.archive.org/web/20150420152310/https://www.cs.nmsu.edu/~rvinyard/itanium/predication.htm">"Predication"</a>. <i>cs.nmsu.edu</i>. Archived from <a rel="nofollow" class="external text" href="https://www.cs.nmsu.edu/~rvinyard/itanium/predication.htm">the original</a> on 20 Apr 2015<span class="reference-accessdate">. Retrieved <span class="nowrap">2014-04-22</span></span>.</cite></span>
</li>
<li id="cite_note-2"><span class="mw-cite-backlink"><b><a href="#cite_ref-2">^</a></b></span> <span class="reference-text"><cite id="CITEREFMahlkeHankMcCormickAugust1995" class="citation conference cs1">Mahlke, Scott A.; Hank, Richard E.; McCormick, James E.; August, David I.; Hwn, Wen-mei W. (1995). <i>A Comparison of Full and Partial Predicated Execution Support for ILP Processors</i>. The 22nd International Symposium on Computer Architecture, 22–24 June 1995. <a href="CiteSeerX_(identifier)" class="mw-redirect" title="CiteSeerX (identifier)">CiteSeerX</a> <span class="id-lock-free" title="Freely accessible"><a rel="nofollow" class="external text" href="https://citeseerx.ist.psu.edu/viewdoc/summary?doi=10.1.1.19.3187">10.1.1.19.3187</a></span>. <a href="Doi_(identifier)" class="mw-redirect" title="Doi (identifier)">doi</a>:<a rel="nofollow" class="external text" href="https://doi.org/10.1145%2F223982.225965">10.1145/223982.225965</a>. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>0-89791-698-0</bdi>.</cite></span>
</li>
<li id="cite_note-Fisher04-3"><span class="mw-cite-backlink">^ <a href="#cite_ref-Fisher04_3-0"><sup><i><b>a</b></i></sup></a> <a href="#cite_ref-Fisher04_3-1"><sup><i><b>b</b></i></sup></a></span> <span class="reference-text"><cite id="CITEREFFisherFaraboschiYoung2004" class="citation book cs1">Fisher, Joseph A.; Faraboschi, Paolo; Young, Cliff (2004). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=DbOa0p1wRD4C&pg=PA172">"4.5.2 Predication § Predication in the Embedded Domain"</a>. <i>Embedded Computing — A VLIW Approach to Architecture, Compilers, and Tools</i>. Elsevier. p. 172. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>9780080477541</bdi>.</cite></span>
</li>
<li id="cite_note-4"><span class="mw-cite-backlink"><b><a href="#cite_ref-4">^</a></b></span> <span class="reference-text"><cite id="CITEREFCordes" class="citation web cs1">Cordes, Peter. <a rel="nofollow" class="external text" href="https://stackoverflow.com/a/50960323">"assembly - How does Out of Order execution work with conditional instructions, Ex: CMOVcc in Intel or ADDNE (Add not equal) in ARM"</a>. <i>Stack Overflow</i>. <q>Unlike with control dependencies (branches), they don't predict or speculate what the flags will be, so a cmovcc instead of a jcc can create a loop-carried dependency chain and end up being worse than a predictable branch. <a rel="nofollow" class="external text" href="https://stackoverflow.com/questions/50959808">gcc optimization flag -O3 makes code slower than -O2</a> is an example of that.</q></cite> <span class="cs1-visible-error citation-comment"><code class="cs1-code">{{cite web}}</code>: </span><span class="cs1-visible-error citation-comment">External link in <code class="cs1-code"><code class="cs1-code">|quote=</code></code> (help)</span></span>
</li>
<li id="cite_note-5"><span class="mw-cite-backlink"><b><a href="#cite_ref-5">^</a></b></span> <span class="reference-text"><a rel="nofollow" class="external free" href="https://gpgpuarch.org/en/basic/simt/">https://gpgpuarch.org/en/basic/simt/</a></span>
</li>
</ol></div></div>
<div class="mw-heading mw-heading2"><h2 id="Further_reading">Further reading</h2></div>
<ul><li><cite id="CITEREFClements2013" class="citation book cs1">Clements, Alan (2013). <a rel="nofollow" class="external text" href="https://books.google.com/books?id=ySILAAAAQBAJ&pg=PA532">"8.3.7 Predication"</a>. <i>Computer Organization & Architecture: Themes and Variations</i>. Cengage Learning. pp. <span class="nowrap">532–</span>9. <a href="ISBN_(identifier)" class="mw-redirect" title="ISBN (identifier)">ISBN</a> <bdi>978-1-285-41542-0</bdi>.</cite></li></ul></div><!--htdig_noindex--><div><div class="zim-footer">
This article is issued from <a class="external text" title="Last edited on 2025-08-07" href="https://en.wikipedia.org/wiki/?title=Predication_(computer_architecture)&oldid=1304642556">Wikipedia</a>. The text is available under <a class="external text" href="https://creativecommons.org/licenses/by-sa/4.0/deed.en">Creative Commons Attribution-Share Alike 4.0</a> unless otherwise noted. Additional terms may apply for the media files.
</div>
</div><!--/htdig_noindex--></div>
</div>
</main>
</div>
</div>
</div>
</body></html>